Papers with automatic annotation pipeline

3 papers
DNASpeech: A Contextualized and Situated Text-to-Speech Dataset with Dialogues, Narratives and Actions (2025.acl-long)

Copied to clipboard

Challenge: Existing TTS datasets lack situated descriptive prompts aligned with speech data.
Approach: They propose a contextualized and situated text-to-speech task to promote more accurate and customized speech generation using DNA prompts.
Outcome: The proposed task promotes more accurate and customized speech generation using DNA prompts.
CogniBench: A Legal-inspired Framework and Dataset for Assessing Cognitive Faithfulness of Large Language Models (2025.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on “factual statements” that rephrase source materials, but ignore “cognitive statements” . evaluating and detecting "faithfulness hallucinations" remains challenging .
Approach: They propose a framework to assess faithfulness of cognitive statements and introduce a dataset to scale easily across models.
Outcome: The proposed framework assesses faithfulness of cognitive statements and scales easily across models.
Concept-pedia: a Wide-coverage Semantically-annotated Multimodal Dataset (2025.emnlp-main)

Copied to clipboard

Challenge: Current evaluations for Vision-language Models remain heavily anchored to ImageNet .
Approach: They propose a large-scale semantically-annotated multimodal resource that extends the range of visual concepts, including diverse abstract categories.
Outcome: The proposed model expands the range of visual concepts, including diverse abstract categories.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations